Papers with spatial grounding

6 papers
GeoGround: Uncertainty-Weighted Multi-Task Learning for Geo-Alignment and Address Defect Detection (2026.acl-industry)

Copied to clipboard

Challenge: Address intelligence in e-commerce requires precise geocoding and proactive defect detection under strict sub-50 ms latency constraints.
Approach: They propose a multi-task learning framework that jointly models coordinate grounding and address defect detection.
Outcome: The proposed model achieves 5.86 gains in address defect detection precision and 4.86 improvements in location prediction accuracy over strong encoder baselines while remaining 75 more efficient than decoder LLMs such as Qwen2-1.5B.
Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamic Audio-Visual Scenarios (2023.findings-emnlp)

Copied to clipboard

Challenge: Audio-visual question answering requires multistep spatio-temporal reasoning over multimodal contexts.
Approach: They propose a new target-aware joint spatio-temporal grounding network for audio-visual question answering . the proposed system integrates audio-vision fusion and question-awful temporal grounding into one module .
Outcome: The proposed method over existing state-of-the-art methods is effective over existing methods . it can focus on audio-visual cues relevant to the query subject by utilizing explicit semantics from the question .
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse (2025.findings-acl)

Copied to clipboard

Challenge: Visual Language Action models have shown promise in decision-making tasks, but have been neglected in previous work .
Approach: They propose a new paradigm for visual language action models that enhances the foundation model prior to action-specific tuning by first post-training it on a curated set of visual and linguistic tasks using self-supervised learning.
Outcome: The proposed model outperforms the best agent baseline on a diverse set of atomic tasks and surpasses imitation learning-based policies in Minecraft.
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation (2026.acl-long)

Copied to clipboard

Challenge: Existing models rely on a single segmentation token whose hidden state implicitly encodes both semantic reasoning and spatial localization . Existing methods rely only on SEG>, which encodes semantic reasoning, limiting the model's ability to explicitly disentangle what to segment from where to segment.
Approach: They propose a method which reformulates reasoning segmentation as a structured conditional generation process over image tokens conditioned on language grounded query banks.
Outcome: The proposed model bridges token-level predictions and pixel-level supervision by decoupling spatial grounding from semantic reasoning through structured language grounded query banks.
Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making (2025.findings-emnlp)

Copied to clipboard

Challenge: Modern embodied AI uses multimodal large language models as policy models, predicting actions from final-layer hidden states.
Approach: They propose a hierarchical action probing method that aggregates representations from all layers, mirroring the brain's multi-level organization.
Outcome: Experiments show that hierarchical probing improves on last-layer embodied models and achieves a 46.6% success rate and a 62.5% gain in spatial reasoning tasks.
Seeing Culture: A Benchmark for Visual Reasoning and Grounding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal vision-language models (VLMs) have made significant progress in cultural understanding tasks . but these datasets often fall short of providing cultural reasoning while underrepresenting many cultures.
Approach: They propose a Seeing Culture Benchmark that requires VLMs to reason on culturally rich images in two stages.
Outcome: The proposed approach requires VLMs to reason on culturally rich images in two stages . the Seeing Culture Benchmark identifies cultural reasoning shortcomings in multimodal models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations